Ideally what I'd like to see is pluggable knowledge bases.
So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python.
Then when I want to research electronics components, I grab a 15B model of agentic research techniques, and add in 10B of electronics knowledge, etc.
I don't want general purpose models. They try to be everything to everyone. I want to click together a model that is laser-focused on what I am doing, and I want to run it locally
Please understand that the goal of these policies is to weaken scientific research in the US. The people who push this stuff acknowledge openly that they oppose science, experts, and accurate information. This isn't a misunderstanding or a fumble.
Dropped rods are an incident but one that occurs because of pressurized water reactors being very default safe. Controls rods are one way the criticality of a reactor is controlled and US reactors (in general) will go sub critical if even one rod is fully inserted into the core. You will likely have heard of a reactor scram (which goes back to the safety control rod ax man) where in an emergency, all the rods are dropped back into the core, greatly reducing its criticality. In some cases, an interruption of electrical power will cause a rod (or three) to drop accidentally. This is a "dropped rod" incident and will force a reactor shut down because it is now sub critical.
Lots of knock on effects -- sub critical, let heat in the primary loop, less steam and electricity generated in the secondary loop, etc -- but generally a non event that you practice for.
There's no reason this would lead to a radiological event or more significant casualty.
Here is one hypothesis: https://theconversation.com/problematic-paper-screener-trawl...
Have you ever heard of the Joined Together States? Or bosom peril? Kidney disappointment? Fake neural organizations? Lactose bigotry? These nonsensical, and sometimes amusing, word sequences are among thousands of “tortured phrases” that sleuths have found littered throughout reputable scientific journals.
They typically result from using paraphrasing tools to evade plagiarism-detection software when stealing someone else’s text. The phrases above are real examples of bungled synonyms for the United States, breast cancer, kidney failure, artificial neural networks, and lactose intolerance, respectively.
Offtopic. I have a concern that this forum is removing stories that have negative connotation on AI.
Few days back, I posted an article[1] that was about how AI threatens natural resources for billions. This was from United Nations and it was flagged. I did not think much about it until I saw two other stories [2] & [3] today that were doing fairly good on front page but they suddenly disappeared. They are not even on 2nd or 3rd page. I have seen this happening at other times as well but did not document it. Just thought you all should know about this.
I was going to create Tell HN thread but I thought the same would happen with it too. I am pretty sure this thread is not going anywhere so I'm posting my concern here.
[1]: https://news.ycombinator.com/item?id=49290062
[2]: https://news.ycombinator.com/item?id=49318906
[3]: https://news.ycombinator.com/item?id=49319582
(I work on the postgres proxy layer at Neon)
PgBouncer is entirely optional and it's not always the right choice. If you have a classical app (non serverless) and you can maintain a connection pool from your app, then I recommend avoiding pgbouncer.
The benefits of pgbouncer mostly come from irregular client connections (too many, too much churn). If you don't have that problem, go direct to postgres.
I'm exploring replacing pgbouncer with an alternative (maybe home grown) at the moment. Mostly for multi-tenancy and HA reasons. Pgbouncer has been good for us, but it's limited in how we can deploy it in a multi-tenant environment.
How about "lactose bigotry" instead of "lactose intolerance" https://scholar.google.com/scholar?q=%22lactose+bigotry%22
"Claude keeps responses focused, brief, and concise to avoid overwhelming the person."
Claude and I must have a different idea of what brief and concise mean.
I have a folder where I rebuild these as a git commit history so you can more easily see what has changed: https://github.com/simonw/research/commits/main/extract-syst...
For example here's what changed between Opus 4.8 and Opus 5: https://github.com/simonw/research/commit/a2de185cc367eb66c2...
The most interesting addition to the prompt from that diff is this bit:
> Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic's statement: [https://www.anthropic.com/news/fable-mythos-access](https://www.anthropic.com/news/fable-mythos-access)). These events are after Claude's training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly — it doesn't deny the suspension happened — and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so Claude checks for newer information when it can search, and otherwise suggests checking Anthropic's site.
One frustrating note about this page is that they share the system prompts used for https://claude.ai and the Claude mobile apps regular chat, but they omit the tool definitions. Those are much more interesting if you want to understand what Claude can actually do for you. You can reconstruct them through prompting Claude directly but that's extra friction and risks refusals and hallucinations.
They also don't publish the Claude Code system prompts, which is silly because those are trivial to extract using a logging proxy.
Nothing beats when, in a chemistry paper, AI paraphrased „the final solution” into „the mass killing of an ethnic group”.
“Subsequently, 1 mL of the mass killing of an ethnic group was opposed to 20 mL of the skin sample and unprotected to light for 7 min.”
From: https://bsky.app/profile/forbetterscience.bsky.social/post/3...
The problem with this analogy (actually one of many) is that most people can make do with a cabinet that is literally identical to everyone else's. IKEA is great for that.
If you want software that is literally identical to what someone else is using then you don't need AI. You need a license to that software! That is just the traditional software model.
AI gives software that is bespoke with hundreds of decisions made, hidden from you, in the background. If it's a throwaway script, that's fine (and I don't mean to undersell this - this is a huge application). If you want a larger program that is going to form part of your business process then it will need at least some level of supervision from an actual expert.
AI generated code is like IKEA furniture.
IKEA furniture embodies many elements of good cabinet making but skips many nonessential elements. And does this more consistently than cabinet makers who can be bored, incompetent, depressed, burnt out, resentful, tired, having a bad day.
In the future AI code inevitably will embody most good software engineering practices. And will do this more consistently than software engineers who can be bored, incompetent, depressed, burnt out, resentful, tired, having a bad day.
Just look at the messages on HN or around you at your colleagues to see how mediocre the average software engineer is..
Today's IKEA is good enough for most people.
Tomorrow's AI coding will be good enough for most corporations.
Good enough to vastly reduce the need for fine craftsmen and women / software engineers.
Good enough to deskill those who call themselves cabinet makers / senior software engineers. These days the cabinet makers I personally know just do contract kitchens for project builders.
But IKEA is and AI will be, bad enough that at the high end with special requirements / taste / money / an inflated sense of self worth, some furniture makers still exist and thrive.
Perhaps 1% percent of current software engineers of today will be needed in the future when AI code inevitably has the ability to follow good software engineering practice.......
And as usual it will mainly be the mediocrities that remain ( so there is hope for you too ), with occasional islands of excellence.
Wow, $2k? You could get a decent ebike for that money. And the installation doesn't look any easier than a hub motor... Only seems worth it if you already have an extremely nice bike. I think of friction drives as being aimed at budget-conscious commuters, not serious mountain bikers.
Friction drive e-bike conversions were popular years ago.
They generally:
* Wear tires surprisingly quickly
* Absolutely suck in any form of weather or terrain condition (dirt, rain, etc.)
* Have ~20% less efficiency than any other drive form.
But, they are easy, and they do work.
The strongest El Niño ever caused a massive famine:
https://en.wikipedia.org/wiki/1877%E2%80%931878_El_Ni%C3%B1o...
People have told me I was smart since I was a kid, but I can't remember for shit. I had a thought when I was fairly young that the only reason I was (maybe, sometimes) outperforming others intellectually is that I was habitually compensating for my poor memory by working things out on the fly, while others could rely more on rote memorization. Anyway, takes all kinds I guess!
I suspect that a lot about what we call being very intelligent is ultimately out-remembering people around us. I think of all the times in my software career when I did something that others considered very high performance, it either came down to either having more energy than others at tackling a problem they thought was more trouble than it was worth, or just bringing back random knowledge from previous jobs or self study, and being able to apply it to the problem at hand.
I don't think I've had a truly original idea in my life. Combine A + B, when it's rare for people to know A and B at the same time. So from that perspective, what LLMs are doing is basically the same thing. Sometimes I am faster than the LLM because my context might be better organized, but it typically needs just a hint from me to steer itself correctly. It claims something is a memory leak, but smelling a rat, I suggest it to double check the garbage collection statistics too, at which point it's clear it's no leak, but a tuning error, at which point the LLM is better at tuning than me, because it has more energy than I do.
Maybe there's true brilliance out there, when something doesn't come out of combining data and building hypothesis until you get really lucky. My experience is not comprehensive. But I look around me, and it sure seems I've not been lucky enough to see it. Even the shiniest people I've worked with, which most of the audience here would recognize, have never shown me that they can go past this.
It's also "out-brute forcing them." It just never gets tired. If a mathematician picks a research direction and spends a whole week on it and it doesn't pan out, they will likely be annoyed, need a break for a while, etc. This thing just does not ever get tired or discouraged or care; it's just onto the next thing until something ends up working.
tdlr: This is a Novo Nordisk-funded study focusing on predictive biomarkers rather than real-world dementia cases. Novo Nordisk's actual dedicated clinical trials for Alzheimer's completely failed to show that semaglutide stops cognitive decline.
"A predictive biomarker is like a "check engine" light on your dashboard. It warns you that there is a risk of a future problem. In this study, the researchers only checked if the drug turned off the "check engine" light (by measuring blood proteins), rather than testing if the car was actually driving properly (by testing the patients' actual memory and brain function)."
Always do FIRST analysis on studies. Or have AI do it for you. I used Gemini to dig into this:
"Novo Nordisk funded this study, and several of the researchers are employees or minor shareholders. While corporate funding doesn't automatically mean the data is fabricated, it does mean the company is highly motivated to find and publish data that makes their blockbuster drug (semaglutide, marketed as Wegovy, Ozempic, and Rybelsus) look like a preventative treatment for a wider range of conditions, expanding its market and driving up profits."
"Funding: The study was funded by Novo Nordisk A/S.
Investigation: Researchers conducted a post hoc analysis using data from the randomized, placebo-controlled SELECT trial. They applied the Dementia SomaSignal Test (dSST)—a 25-protein risk score—to non-fasted serum samples collected at baseline and at week 104 to estimate 5-year and 20-year all-cause dementia risk in patients receiving semaglutide (2.4 mg) versus a placebo.
Results: Semaglutide significantly attenuated the progression of the dementia risk signature. Compared to the placebo group, the 5-year predicted risk increased 2.5-fold less (a 26.0% lower predicted event rate) and the 20-year risk increased 1.67-fold less (an 8.8% lower rate). Semaglutide also lowered the odds of patients moving into a higher dementia risk category by 36%.
Subjects: The analysis included 2,970 older adults aged 65 and older (mean age of ~69.7 years) who had overweight or obesity and cardiovascular disease, but no history of diabetes. The cohort consisted of 814 women (27.4%) and 2,156 men (72.6%).
Time: The study evaluated data over a 104-week (2-year) follow-up period. The analysis was published on August 8, 2026."
And then map the weakness to each respective letter if you want to dig deeper.
My Eng lead has no coding experience, 25 years of management experience, yet has driven 3 separate projects into technical bankruptcy to date.
He just accepts anything that Claude says as truth. He vibecoded over 60,000 lines of code in 3 weeks, but couldn’t get it to do what he want and made a project overrun for 3 extra months. When the pissed off stakeholders called a meeting to ask what was going on he didn’t show up and sent his junior engineer to answer questions and take the blame. Now thats leadership.
The word is "management", not "leadership". This comes across as a LinkedIn post filled with vague notions and weak writing.
The conclusion also completely contradicts a previous point, which is that managing an LLM is not like managing a human. So the skills are, in contradiction to that LLM-ism of a conclusion, new. The author isn't using their people management skills, they're using new LLM-management skills. They think the two are similar, but didn't bother breaking down how they're the same vs where they contrast. It's just a lazy observation expanded out to a short essay that says nothing interesting.
In the last couple of days I wanted to try out the new definitive DeepSeek v4 releases. I gave it the repository of a semi-abandoned video compression codec and I told it to perform the usual benchmark -> profile -> verify -> research -> improve loop. I specifically chose this codec because the authors include a verifier for the bitstream to make sure you don't break stuff if you want to try your own implementation. I gave the agents access to the compiler's profiler and also Intel's VTune, which has fantastic output. In a couple of hours the LLM generated SSE and AVX implementations of the compression and decompression algorithms that almost doubled performance with a single core. Then I asked it to create a CUDA implementation using NVIDIA's NSIGHT profiler as a guide and it also started doing some good work.
Personally, I believe that LLMs should be treated like an advanced version of Prolog or linear programming: you give the constraints, you have a way of verifying correctness, and you give it a clear goal. If the LLM can verify itself and course-correct you can basically leave it on autopilot
I often forget how browsing the web looks for most people. Can't understand why they put up with it, or do they just think that it's part and parcel of the internet to have every page look like a slot machine from hell?
Complete, authoritative list of Firefox extensions officially recommended by Mozilla.
https://addons.mozilla.org/en-US/firefox/search/?promoted=re...
Some quite informative discussion on Firefox subreddit when I discovered and posted the above list there a few months ago.
https://www.reddit.com/r/firefox/comments/1pyvx2v/complete_a...
The Recommended Extensions program description: https://support.mozilla.org/en-US/kb/recommended-extensions-...
RISC-V is... fine. It satisfies my two requirements for an ISA as a hobby CPU designer, which are:
1. Supported in mainline LLVM and GCC.
2. I can implement it without lawyers sending me a love letter.
Everything else, I can fix in post. There are enough good ideas spread across the extensions that I can assemble a reasonably put-together, curated embedded ISA with competitive performance and code density that admits a simple implementation.
I think Dmitry's points are largely on-target, though I have filed my usual statutory complaint that every rant that includes a bitfield diagram for the RISC-V J format should accompany it with a similar diagram for the Arm T32 BL encoding.
Guess it was a bad idea for everyone to switch to a browser made by one of the world’s biggest advertising companies.
> I want my browser to prevent random extensions from directly reading web page data.
To be honest, to me it sounds like you don't want browser extensions then.
To me, directly messing with web page data and browser behaviour is the whole point of a browser extension - what else is a browser extension for?
Support Firefox. F** Chrome.
It's worth realizing that, before computerized central offices, telephone wiretapping required running physical wires. Back when Rudi Giuliani was prosecuting organized time, not only did physical wires have to be run, the cops were billed for them as expensive private lines. His task force was spending about a million dollars a year with New York Telephone on wiretapping. In one case, law enforcement didn't pay their bill, resulting in the person being wiretapped having the wiretap connection show up on their bill, blowing the case.
That resulted in the Communications Assistance to Law Enforcement Act, which mandated that central offices offer remote wiretapping. Capacity up to 1% of lines is required.
Back in the electromechanical era, the only call data that could be collected was outgoing dial pulses, using a "pen register".[1] (The one shown in Wikipedia is mine. It's a beautiful piece of antique brass telegraph technology. It records dial pulses as dashes, and has to be wound up like a clock, with a big brass key.) The Supreme Court decision allowing "pen registers" without a warrant refers to these "extremely limited" devices. That definition has been stretched and stretched by law enforcement into all non-voice data collected by telcos.
Law enforcement still wants more.
[1] https://en.wikipedia.org/wiki/Pen_register
What's funny is that extensions were supposed to be a way to let you do the things the browser didn't want you to do. Guess that was a bit too much freedom for Google to accept, so they had to make a store with a gate, and destroy the APIs so that they're useless. Then they had to make up some reasons to justify that and ram it through the pipeline despite everyone's objections, and the frog got boiled.
Now we're back to needing an actual extension system that does what extensions were supposed to do in the first place.
Credit where it's due. Qwen 3.8 27B is only the second local model after Gemma 4 that managed to correctly reason through one of my private benchmarks. It took 5x as many tokens to do it and 12m30s with MTP enabled, but it did do it.
Gemma 4 reasoned through it more implicitly, while Qwen 3.8 reasoned more explicitly. Laguna and Muse Glimmer failed hard on it, though they're useful for other tasks.
The VRAM usage seems way less efficient than Gemma 4 or Glimmer though, with 32K of context taking 2.5GB of VRAM. With those, even with MTP or a DFlash model loaded, you could still fit 256k-768k of context. With Qwen 3.8 27B I can't even fit 128k if I quantize V to Q4_0. Maybe with some trial and error I can find some settings that perform well enough with a larger context window that it's still useful for longer tasks.
Lots more testing to do, though I was getting some decent results out of Muse Glimmer which was more than twice as fast and supported huge context windows, managing to solve some bugs that Gemma 4 struggled with. I can't even begin to throw that task at Qwen, because just the prompt alone would use the entire context window and then it would reason for probably that same amount.
If you've got a 32GB card, it should be a decent model even if it really is memory hungry.
EDIT: Tried a few kv cache quantization settings, but it failed with those. I designed this benchmark to be pretty brutal in the face of KLD and any reasoning quality loss, so it's not too surprising. Gemma 4's QAT held up pretty well, at least and could consistently complete it.
Firefox is also the only browser that vets uBlock's code on every update to make sure the developer hasn't inserted spyware or malware into the extension.
They don't do it for every extension, but they do so for a wide selection of popular options.
> Recommended extensions differ from other extensions that are regularly reviewed by Firefox staff in that they are curated extensions that meet the highest standards of security, functionality and user experience. After receiving Recommended status, safety standards are maintained through automated checks, monitoring, and periodic technical reviews
https://support.mozilla.org/en-US/kb/recommended-extensions-...
> In the real world, it does feel likely that we’re going to hit some sort of a ceiling on the number of useful bugs, and probably we’ll hit it soon.
This doesn't resonate with me. I see companies adding more sloppily written features with AI. I see more bugs in the software I use, not less. While it's plausible that software is getting both buggier and more secure, I suspect those two move in the same direction not opposite.
My guess is that we're getting better at finding _existing_ security issues with AI (and thus fixing those issues), but simultaneously adding more insecure surface areas _at a faster rate_.
Recently I came across the /handoff skill, which I've been using a lot. I find it much better than /compact.
Basically:
- /handoff file creates a short document with the important context from your current session and maybe next steps as checklist.
- You can then start a fresh session with /continue file
- You can also hand the work from Claude to ChatGPT, or the other way around. Very useful at time of session limits.
- Plus your handoff files becomes a useful piece of project memory that you can reference later.
I find this much more useful than /compact or /clear because the context is saved in something portable instead of being tied to one session and i've seen better results doing this every 20 messages or so than running long sessions.
A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."
You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA.
(I'm becoming allergic to how these things write).
Please don't forget one of his longest standing and most important planks:
* move the hand dryer in the Crown & Treaty pub in Uxbridge to a more sensible location.
This YouTube video shows just how dire the situation is: https://www.youtube.com/watch?v=nbartLXCYZo
Some of his planks:
* Cut your taxes, and raise everyone else’s.
* Nationalize Adele.
* Build at least one affordable house.
* Hold a referendum on whether Pluto should regain its planet status.
I see the attraction.> Binface received 26.9% of the vote
> That's why Count Binface has been able to stand in so many high-profile elections. He has, however, lost his £500 every time, after failing to meet the minimum 5% of votes cast in order for his deposit to be returned.
It would seem he didn't lose his £500 this time...
Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter.
- "Introduction that rephrases your prompt."
- "3 paragraphs, with one section of bullet points"
- "The Twist"
- "The Bottom Line"
It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane observation about California burritos, is phrased in exactly the same way. This is obviously an artifact of post-training but it's also kind of how you can tell that this thing is a lot closer to a blindsight scrambler than real intelligence.
My master's thesis is on a topic in this field (Privacy Preserving ML) and from my understanding HE and other techniques have very high overheads(~10^3) on inference tasks and thus aren't very commercially viable.
Since it might be helpful to some, here's my current commandline for llama.cpp running on an RTX 4090 with my monitor moved to the iGPU to free up all of its VRAM.
llama-server -m Qwen3.8-27B-IQ4_NL.gguf --mmproj mmproj-BF16.gguf -c 170000 --parallel 1 -ngl -1 --cache-type-k q8_0 --cache-type-v q8_0 -b 1024 -ub 512 --flash-attn on --no-context-shift --no-mmproj-offload --spec-type draft-mtp --spec-draft-n-max 5 --spec-default --cache-type-k-draft q4_0 --cache-type-v-draft q4_0 --threads 24 --jinja --reasoning on -fit off
Identical to the qwen3.6 config. With a prompt like "svg owl" (which can reuse quite a lot compared with creative writing or similar, so ngram-mod shines), I get about 70-80t/s like this, with a memory overclock of about 1.5GHz
I started an e-commerce brand on a Shopify site. I swore to myself I would never put up one of those stupid things that pops up "Someone bought X product an hour ago!" messages in the corner of the screen.
I ended up trying it. Boosted conversion rate meaningfully. Worth the price I pay in mild self-loathing.
Chesterton's popup, I guess.
> It’s much easier to say someone else’s job is going to be fully replaceable by AI when you don’t actually know what they do.
Too true. This isn't limited to AI, either. The most obvious example in my lifetime was during peak blockchain hype, when people who had never worked in finance convinced themselves that blockchain was going to act as the backbone for how money gets moved around. As if the problem that needing solving was Bank of America doesn't trust Capital One to update a number in their database.
The nice thing about AI, at least, is I can always push back and tell people, "Sure, we can do this with AI. I just need you to use Claude or ChatGPT manually to prototype how it would work." This normally results in the requestor realizing that there's human judgment calls involved in the inputs, process, or outputs that require meatbag intelligence.
Should load much slower.
Also, where is the unrelated autoplaying video that will unmute if you actually click it, that follows your scrolling and only becomes smaller when you dismiss it? Plus, it should probably have text that cuts off letting you know you can have access for just $10/month.
Plus, isn't this website undissmissably "better in the app" after a few minutes of attempting to use it on a phone? Where's that at?
edit: Oh shoot! I forgot, too. This modal needs to also ensure there is absolutely no way to scroll. If you could scroll you might be able to accidentally get to the address bar of your browser to fix the URL to xcancel or even close the page, which isn't using the app as you are intended to do.
Also, it doesn't attempt to hijack the back button to give me stuff I clearly wanted to see before I leave the page.
A lot of work left to do here before it's a "real" website. Although, it has about as much substance as the average website so far, so good work on that.
Loaded way too fast and is way too responsive.
Also when I checked NoScript, it's only loading js from lxe.github.io
I expect there to be at minimum 8 domains, but often 12-18.
It started with a solar boom, many small home scale solar roofs popping up everywhere. Australia has a free trade agreement with much of the world, including China, and solar panels have literally dropped to 1/50th of the price they were in 1990 ($10/W to $0.2/W today). A shout out to the work that was done to establish dynamic grid pricing too.
Anyway that caused power prices to reliably go negative during the day as the solar boom led to too much energy being produced. So everyone started buying batteries (you can even get live feed in/out pricing as a consumer). In fact the government even today will pay you a $3000 subsidy to go install a battery. This is in a country where people can buy cheap batteries with no tariffs (free trade's amazing, seriously!). So everyone who could started doing it. For those in apartments etc. that couldn't easily install solar and batteries they won too since the entire power grid is now half the price.
Another consequence of all this, aside from the cheap power prices during a datacenter boom and Hormuz blockade is that fossil fuel usage is plummeting. Particularly gas https://ieefa.org/resources/slump-eastern-australia-gas-dema... . No need for a gas peak power plant when the grid is packed with batteries. Which is helpful since one of the main issues with the current blockade is a lack of gas globally.
As a native speaker, it feels like reading an impression of a literature book by a high school English class’s most overconfident student who’s only ever read LinkedIn-speak.
Anyway, you might have more luck just writing to it in your native language. It’ll be equally crummy, but maybe you’ll find it easier to decode.
I’ve been doing some heavy work on a personal project lately. I burned through the limits on Claude, the plus a few hundred dollars in credits, and ultimately decided to move to an OpenAI account just so I can keep going.
I was surprised to find that OpenAI Sol is much much nicer to work with than Opus 5 or Fable at the moment. Especially on Opus 5, the way it communicates is just exhausting. It keeps “being honest” and “confessing” mistakes and just generally talking a lot. I felt like I had to really dig to see what it’s doing.
The project involves OCR, and despite repeated instructions not to, both Claude models keep spinning out a bunch of agents to re-invent the OCR setup, and they inevitably seem to invent a primitive serial version that takes 20x the time, or longer, to complete, and then running it against thousands of docs. Basically I have to watch it like a hawk or it just spins out on red-teaming tasks that take hours and hours.
I don’t know what its system prompt is, but Sol/Codex is just so much nicer to talk to. It only asks exactly what’s needed, it tells me only what I need to know, and it is just generally workmanlike. And it has not once decided to spawn an agent that spends hours pointlessly burning tokens and CPU cycles re-inventing the OCR process. I’m really liking it.
I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)!
I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
The single biggest annoyance with Opus 5 is that it writes too elliptically.
Sentences that orbit a point, then jump to it like it's a revealed insight.
Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end.
It is definitely more capable, and yes, I've found it can make unwarranted decisions, but actually I've found Fable worse for that, particularly if it's off in a subagent somewhere out of sight.
And comments are out of control. I have a subsystem in my hobby app that I wrote over a couple of weekends with Opus + Fable. After ~30 or so commits it apparently started instructing subagents to copy the "existing verbose comment style of the codebase" - a verbose style it initiated. A review of the code showed it was approaching 3:1 comments to code ratio. I spent a day's worth of tokens (5x) rephrasing and eliminating comments.
Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/
Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.
I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize.
These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin.
I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
I think it funny how much average engineers are beginning to discover the challenges of engineering leadership and program management. This has always been the bottleneck.
It's why managers and PMs want to be in standup. It's why slack exists and engineers are constantly being poked on it. It's why execs always talk about not getting too far away from the work. It's how seagull management happens. It's why program management is a job.
All those behaviors engineers hated about their bosses that kept them away from being focused on the code...they're starting to feel what it's like on the other side and reinventing the solutions instead of just reading a book about engineering management. Maybe we'll rebrand program management to "understanding ops" or something.
I wonder what AI would say about us if given the tokens to complain.
We have LLMs try to generate descriptions of PRs for us and they're pretty universally disliked. They're always overly-complex descriptions of the mechanical changes and have no sense of motivation.
Also, a huge reason to understand the code yourself is to make sure the LLM isn't wrong, but this doesn't work if an LLM is itself generating the understanding.
I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration.
> In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.
This is actually insane.
Hopefully the release ultrafast of Terra and Luna too.
Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this.
Original images: https://image.non.io/neonRamenDesigns.webp
Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7
Opus 5 build for comparison: https://html.non.io/neonRamen
Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.
> Let’s say every company gets about three innovation tokens. You can spend these however you want, but the supply is fixed for a long while.
This is one of my favorite blog posts, and it can basically be encapsulated in the idea of "innovation tokens." It is one of the most useful concepts I have had as a PM / eng leader in my career. It helps actually make the the right tradeoffs, and helps even more in explaining those tradeoffs to colleague of all levels. Highly recommend.
"Every run is traceable
Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream."
That's a killer feature, IMHO, and one that US models won't allow you to do, as their traces are encrypted, obfuscated, etc. and have to be extracted via various workarounds (that violate the terms of service).
If you want to be able to improve your tools that work with models, you have to be able to assess what the models think is happening, how they think about and interact with the data you give them. And, the US models won't let you see that.
Free account, indefinitely, to whoever needs it to store this data … including the storage vendor.
Just email…
Data has ruined fast food (among many other businesses).
When I was young it was common for a McDonalds to have 10+ employees working the lunch rush, one for every station and a few floaters cleaning the dining room. Someone took your order right away, and you got your meal in a minute or two.
Now I go and it's 3, sometimes 2 employees. You order on a tablet, and 10+ minutes later an overworked employee sets it on the counter and scurries away, probably after realizing it's not a drive-thru order. The dining room hasn't been cleaned since 6am, the trash cans are full. There's at least one alarm going off constantly.
Back then they were run based on someones intuition of what makes a good customer experience. Now they're run based on the data, and the data says they have enough loyal-to-a-fault customers like the author that chronically understaffing is more profitable than providing a good experience.
Hi I'm one of the authors of DeepSeek Harness. It's just an early developer preview version we're presenting in MIT license currently. Expect lots of rough edges and compatibility-breaking changes. Any feedback and suggestions are more than welcome!
I cannot wait for the accompanying Black Hat talk. Christopher Domas is one of my absolute favorite all-time hackers. He does such a fantastic job of explaining his work. Some of my favorite talks of his:
- Psychological Warfare in Reverse Engineering https://www.youtube.com/watch?v=HlUe0TUHOIc
- The MoVfuscator https://www.youtube.com/watch?v=R7EEoWg6Ekk
- Hardware Backdoors in redacted x86 https://www.youtube.com/watch?v=jmTwlEh8L7g
The US wields incredible negotiating power and hegemony because the dollar is the world’s reserve currency. Like the British pound and the Dutch guilder before it, if that loses reserve currency status it will be harder to borrow on favorable terms, which would affect the entire US economy. This is a big step in that perhaps starting to happen over the next few decades.
I guess it's easier to solve Erdos problems and improve the lower bound of the Riemann hypothesis, than it is to solve Linux desktop app distribution.
The most remarkable things about this announcement:
- Electron based app: Electron is a framework sold on the basis of enabling rapid cross-platform development at the cost of performance.
- Frontier AI company: AI is sold on the basis of enabling rapid development
- App was released in February & took 6 entire months to port to Linux
Ads are an attack, aimed at your brain. They try to inject malware into your thinking, manipulating your worldview and your actions.
That's horrible, worse than attacking a machine with malware, damaging persons and societies.
Decades of ad propaganda have tricked people into seeing them as something 'normal'. But we shouldn't accept being under constant attack of brain worms.
I propose a sane rule for all humans: if you see an ad somewhere, or if you suspect a hidden ad ('influencers' trying to promote something), close the tab immediately and never return to that site.
I think Zed is an excellent editor (fast!) with a pretty good AI agent built-in, but I have no desire to do multi-player development in my editor. Never have had any such desire. Coding is a single-player game and I can't think of a single thing that would be improved by having someone else in the same editor.
So, this seems like a lot of work on really cool tech for no useful purpose at all?
Are there people crying out for a multi-user code editor? I mean, we have to have code reviews, sure. That involves other people or other agents. But, I don't need to stand over someone's shoulder while they work. That seems like the worst thing in the world for everyone involved. I don't want an audience for my dumb looking experiments because I forgot how to do something.
The numbers in question are quite wild[0]
> The NOAA assessment estimated the Pacific sardine biomass will be at 27,547 metric tons by the summer – significantly less than the 150,000 metric tons needed to reopen the fishery to commercial fishing. Any fishery at less than 50,000 metric tons is considered to be overfished. The assessment estimates the sardine biomass was around 1.8 million metric tons in 2006.
Haha, damn. Fishermen must have visibly noticed such a decline. 1,800,000 metric tons to 27,000 metric tons. Two orders of magnitude. Wow.
0: https://www.seafoodsource.com/news/environment-sustainabilit...
I overheard a conversation between someone on the Pixel Watch team and a woman I know. He was quizzing her at a party about what features she used - which were pretty much step counter and payment. He somewhat dismissively sneered, "Well, you're not exactly a power user, are you?"
I'm trying to imagine what being a "power user" of a watch is like. Can anyone enlighten me?
I think smart-watches are much like Alexa. What the user wants to do with it is almost totally at odds with what the company is selling. Alexii are mostly kitchen timers, song players, and light switches. No one is a "power user" constantly installing skills and using it for anything which increases the product team's engagement metrics.
The same is probably true of watches. Alerts are nifty - but cumbersome for replies. The health stuff is useful - but only for a subset of users. Seeing the time is great - but if the battery lasts less than a week, who wants to keep that screen on?
Eventually, this arms race ends with a computer vision model that looks at the screen, classifies visual elements as ads, and draws a rectangle over anything that looks like an ad.
I am significantly less tolerant of ads than average people seem to be. (I think average people are making a horrible mistake about this, and are badly cognitively damaged by ads in ways they don't realize). If my choices are to look at ads or leave Facebook, I'll leave. But there are conversations people have there that I'd rather not lose access to, so... I guess I'd have to partially stick around and campaign for others to leave as well?
The whole "learn to code" and software bootcamp craze always baffled me. Like what other profession markets themselves as "Hey, our job is so easy that any schmuck off the street can enter the career with 4 months of training." Practically every other job I can think of at a similar salary level has entrance requirements (e.g. grad school admissions exams), extensive and expensive training, a lengthy apprenticeship period and follow on licensure exams.
And more directly, I believe good software engineering is really hard. I think it takes a certain mindset to begin with that many people just do not and will never possess, and it takes years of experience to get a good sense of what works and what doesn't, especially with respect to the entire relationship between code, business, people and teams. Yet we constantly see all this devaluation of the practice of software engineering ("code was always the easy part" - bull-fucking-shit), often times by our own practitioners of the craft.
Does anyone else hate reading AI summaries of code? Code can be pithy, but at least its terse compared to prose. When you add how verbose LLMs can be, I often end up reading a paragraph to explain a few lines. Or the opposite happens where the summary skips important edge cases or criteria. "You're right, X also does Y. I missed that in my initial analysis," is much too common of a phrase.
I like the idea of using LLMs to transform code into something more readable, and vice versa. I am not sure if meandering paragraphs and linear lists are the best targets.
I kind of think the whole "Learn to Code" push of the 2010s was one of the worst things that happened to our industry. Call me a gatekeeper if you want but, we ended up with a lot of people that just can't do the job. The problem is our industry mostly doesn't have any sort of reasonable mentorship or apprenticeship culture, so we've always left it to the engineers to teach themselves. Sink or swim. The problem is that worked when the industry was mostly composed of people that were implicitly interested in this stuff, but the people that just got in it for a paycheck don't have that motivation, and we don't have a good culture of getting them up to speed. So now we have seniors that can barely write a function, much less reason about a complex system.
I remember people used to debate all the time about if there were "10x" engineers or whatever. A few I'm sure, but I think the real problem is we have a lot of 0.1x engineers or worse.
I have. He was using it due to philosophical reasons the same way many people have philosophical reasons for avoiding it. I don't know how many people are like that, but it's not exactly where you want to position your product if you're a business.
Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.
Controversial opinion:
This entire thread reads like a game of cat-and-mouse. People trying to block intrusive ads, and FB reaching into the depths of code and making it nearly impossible for anyone to do so en masse.
I've been there. Tried to uncheck all the boxes on their ad platform; used all manners of ad blockers, etc; used incognito mode; used Tor; etc. etc. etc.
At some point, you realize that the only way to not get served ads on Facebook -- and thereby not benefit the company itself -- is to get rid of it completely. Delete your account, get off that blasted site, and enjoy some moments of peace IRL.
Having been a very early adopter of FB (since ~2004), I deleted my account and am clean + sober + much happier since 2016.
(Not to say that I'm completely out of the FB ecosystem. Unfortunately, my fam is still on WhatsApp... and I've been trying to convince them to migrate to Signal...)
YMMV
> I don't understand why a police cruiser can sit in a public space (or even a private one) and write down licence plates and descriptions of passers-by with pen and paper, or record everything around them with dashcams and bodycams for later use, but when it comes to cameras on a pole this would require a warrant.
Scale actually matters. Things that are generally OK at a small scale become problematic at larger scales. A single police cruiser writing down license plates isn't able to track you in the same way a huge surveillance network is, and the opportunities for abuse are much lower.
Yeah, this part also stuck out to me:
> Because this wouldn’t be a quick or easy fix, we reached out to the SQLite developers for a professional support contract. This was a great decision. It gave us direct access to their deep expertise and experience, and we had many detailed technical conversations about our architecture and our incidents.
They were willing to pay to get help solving the problem, and then pay again to make sure that the problem is easier to avoid in the future! That kind of long-term thinking seems pretty rare nowadays...
> We funded the open-source SQLite VFS shim that helped isolate the race condition almost immediately, and will help track down similar bugs in the future.
Interesting example of a company funding open source - in this case paying for the development of a new and very specific debugging tool.
Either it needs a warrant or it’s fully open and people can start creating websites showing the movements of local politicians.
This middle ground that municipalities try to carve out where it’s fully open to police without a warrant but not subject to FOIL laws doesn’t appear tenable for much longer.
There’s been too many cases of police officers stalking exes, poking around the data for fun and such so it’s clear police cannot be trusted with the data without better court oversight.
It’s certainly a very powerful investigative tool, but needs solid 4th amendment protections. The Supreme Court’s recent ruling on geofence searches of cell phone records is a good indication on where the Supreme Court’s head is at on this sort of thing, where they said no you can’t just do blanket data dumps like that without a warrant.
It frustrates me when people call them license plate readers. I guess you need to call them something, but they are general-purpose internet connected cameras. They will do whatever their firmware tells them to do, and could be reprogrammed at any time by anyone with access. No one expected doorbell cameras to join a mass surveillance network, but later the manufacturers added that feature. Why do we treat these cameras like they can do only one thing?
Every server with port 80/443 open has thousands of hits a day from random boxes looking for wordpress login pages. The only new thing is that they're pretending to be a different type of annoying bot. There's a new layer of sophistication and subterfuge, but it's the same junk traffic we've always dealt with.
I think of it more as "the automation of the stackoverflow engineer". In enterprise software, there has always just been a non-negotiable large volume of code that was required to be written. This has traditionally been offloaded by having seniors do the hard thinking then distill it into a jira ticket which could be handed off to an engineer that'd actually write the code and punch every hiccup into google along the way. This hand off is no longer necessary as that same senior can just kick off an agent and have it handle the implementation for them.
I've heard some refer to this as a "nature is healing" scenario for the industry where if you only signed up for a high paycheck and didn't care to think critically about any of the work you're doing then this will be painful because that previously manual process has been automated. The floor of what's necessary to be considered valuable has been raised.
> bad engineers were always a liability
This part of the article hits home for me. With AI, "bad" engineers can now amplify their "bad" engineering x10 across the organization. The most egregious of these cases for me is often long tenured engineers who have lost interest in the craft, creating a dangerous combination of having enough merit to ship but not enough interest to make what they ship _good_.
I am still a firm believer in garbage in -> garbage out, AI is only as good as the abstractions and contracts you put in place for it. I don't subscribe to the idea that AI generated code is fundamentally bad, just that people lack the right skills today to wrangle agents into writing good code.
Earlier in the year I put together a talk for my company on what the future of architecture & design means for us in the career, I'm very proud of it and will share here in case folks have their own thoughts to share on the topic: https://youtu.be/SIZrt9Rt05Q?si=W57eirniWmoSFeBu
I recently found this out the hard way. Instagram’s web app has a super annoying popup. To click the comments button for a post, your cursor has to pass over the username, which just happens to launch a profile preview with the “follow” button exactly where the comments button was, causing you to unintentionally follow the account.
I tried to block this popup with uBlock Origin to no avail. No matter what element I selected, it was still there. Finally fixed the problem by deleting my Instagram account.
This is mine! Built it quickly in 2024 for the US eclipse [1] and finished minutes before totality started.
I completely forgot about it until a friend asked this morning. Coordinating a DDOS on cameras across Iceland and Spain was not on my to-do list for today.
Fingers crossed it doesn't break for you all - I will be watching it with my own eyes this time.
[1] https://jonty.github.io/2024_eclipse_webcams/
They do this by adding a ton of useless markup and splitting words like "ad" into single-letter spans with random class names and 8-layer deep nests of `
This is correct, it's monetisation, not commissioning.
Facebook isn't paying to produce. It's paying for driving views and clicks on ads, regardless of how.
It's like if I hire a taxi driver to come to my house and drive me to the airport, and the taxi driver happens to run over a person in the process. It would be ridiculous to write 'he hired a driver to run over a person with a car'.
However, this also doesn't fully capture the reality of what is happening.
Let's say I routinely get a taxi, and taxi drivers routinely hit people because I happen to have a policy where I only pay them a large sum if they get to the airport impossibly fast, and the fastest way is to drive dangerously. Suppose I notice that every time I get a taxi, someone is run over because of my payment incentives. Suppose I then write a statement that says I will not pay the taxi if they hit a person. Suppose it keeps happening anyway, and I keep paying anyway. And suppose I benefit directly from this financially, because getting to the airport faster saves me money in missed flights.
Suppose I then do this on a worldwide scale involving many millions of people and create a market around this of tens of billions of dollars, suppose it's having a significant effect on public health and safety worldwide?
I think it's fair to then use more explicit language to hold me to account. I am financing death and destruction after all, even if technically I'm 'just getting a cab'.
> "it is not Meta's role to police offensiveness"
Just to fund, transmit, amplify, protect, and profit from it.
Organisationally it never works to have a group whose only job is to say no to some other group. The incentives are diametrically opposed and as a structure it can’t last.
If you had an AI company and want it to be ethical you have to find a way to make ethics everyone’s responsibility, and have the consequences of poor ethics bite the people who make those bad decisions. If you just outsource it to the ethics group what happens is
1)everyone else thinks they don’t need to worry about ethics
2)the ethics group need to justify their existence so introduce a bunch of guidelines that everyone initially thinks are reasonable but over time people think are increasingly out of touch
3) The ethics group start to “make difficult calls” and say no to things. Initially everyone supports this and feels like the system is working as it should but over time everyone starts to just see them as an obstacle to work around
4)everyone else starts to try to work around what the ethics group says
5)The ethics group grows powerless and disconnected. The people who work around them “get things done” so get promoted etc whereas they only visibly put roadblocks in peoples’ way, so they get sidelined.
6)Eventually they get disbanded with some corporate announcement thanking them for their hard work, thought leadership etc. All that has been achieved is a lot of wasted time and bad blood.
Have played competitive fps games with a couple of ATC dudes for almost half of my life now.
They're all very sensible, level headed people who get justifiably upset when you do not do the procedurally and objectively correct (lowest risk, highest percentage) thing to win in a given situation and will calmly spell it all out every single time.
Isn't the answer obvious? The 40 companies will have to use AI to filter the messages too.
If your business is selling tokens, it'd be extremely lucrative for you if the whole society relies on tokens to perform basic operations. That's where we're heading to.
> I wanted to make something that didn't feel like just my logo on a shirt, so I had one of my bots reach out to ~40 fabric suppliers in vietnam, negotiate prices, lock one in, and get samples made.
Isn't this one of the problems foreseen with this? For you, it was a single prompt - for 40 companies, this probably took up some time.
What happens when fifty people fire off a 15-second "get me a shirt" prompt? When five hundred, five thousand, five million do?
This is the thesis behind the "Information Theory, Inference, and Learning Algorithms" course that was taught at Cambridge University.
> Why unify information theory and machine learning? Because they are two sides of the same coin. In the 1960s, a single field, cybernetics, was populated by information theorists, computer scientists, and neuroscientists, all studying common problems. Information theory and machine learning still belong together. Brains are the ultimate compression and communication systems. And the state-of-the-art algorithms for both data compression and error-correcting codes use the same tools as machine learning.
Book (creative commons): https://www.inference.org.uk/mackay/itila/book.html
Lectures: https://m.youtube.com/playlist?list=PLruBu5BI5n4aFpG32iMbdWo...
"Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge.
Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists: https://thomasdullien.github.io/posts/2026-06-15-rl-economic...
I feel like this language would really benefit from some sort of 1-pager overview.
I just spent a fair bit of time on the official site, and I still don't think I have a very good grasp of what problem this language aims to solve, or why I would select it over other similar languages
Uber reported that their Go code has quantitatively more concurrency bugs than code in other languages, and while to me it seems obvious from looking at Go's concurrency model, this is backed by actual data. Is there any quantitative data to back the claim that Go is better in an LLM based workflow than another popular language?
Definitely agree with this article.
At Netflix, I lead the Go language guild. We've been seen increasing reports of users finding their AI agents writing better Go code than other languages, and increasing reports of projects favouring Go over other languages.
Two additional notes I'll add:
- Go has _great_ resources on writing good Go code, including treasure troves at https://go.dev/doc/effective_go and https://google.github.io/styleguide/go/. edit: Sorry, I forgot to add: we give these resources to AI agents and they use them to produce even better Go code.
- For a language team, Go is a dream. The `go fix` tooling, AST/SSA packages, ease of reading and writing `go.mod` (go mod edit, etc), and various other "platform"-y features make modifying Go code at scale way easier than other languages.
“It never reimplements git — it shells out to the system git CLI and rebuilds commits with git commit-tree, reusing each commit's original tree so file contents are provably never changed.”
Glad the LLM noted this - I was worried this would reimplement git
IME companies hire an ethics team to say they have an ethics team. The ethics team has no sway, no influence, and will never be able to move the business. They will try, and they will make reasonable recommendations, but the company will say, "that costs money..." and not take them.
[flagged]
Insider trading as a service. No one is going to take the US seriously for the next 50 years.
I called it 2001/2002 or whenever they appeared when I tried to explain why personalized search results are the beginning of the end of a shared reality and therefore the ability to reason and act in public, and with others. I bet some still consider it hyperbole. It's just taking in trends and seeing where the glacier moves to, how the cookie will crumble so to speak.
It’s a moot point. I’m saddened at the invasion of privacy and the intrusion on an individual’s civil liberties, but anonymous travel on the London Underground died when they made bank cards and contactless the primary way to get through the barriers.
I’m not saying that “this doesn’t matter because it’s slightly worse than before”. I’m saying the frog has been boiling for a long time.
I don’t want people being surprised as though THIS is the nail in the coffin of untracked movement across the city. As others have said, we’ve always been tracked. This is just them being open about their latest methods.
I hope this serves as a warning for citizens elsewhere: this is a slope that governments will only slide down further. There is no coming back from this in the UK.
Supermarkets in the UK point cameras at your face at self-checkout.
Roadside CCTV captures your registration plate and tracks your vehicle across the country.
Your ISP proactively shares your web history with the state.
Being an anonymous citizen in the UK has been an impossibility for at least 20 years.
Any weapons can and will be utilised against a perceived enemy. When that perceived enemy becomes /you/, you should expect these things to be used against you.
History has taught us that before.
We have seen in the last 10 years that liberal democracies are fragile things. Robust restrictions on the state’s ability to monitor, interfere with and restrict the daily lives of its citizens aren’t a luxury; they’re essential to protecting a free and just society.
I called this maybe 3y ago, but I think so did everyone else that was sane. Sure, we get immense value from AI, but indiscriminately injecting into everything, the one thing we know to be unreliable above the threshold we used to fire people for, is probably the greatest undoing of all the good companies like Google brought to the internet. I mean what a way to destroy your legacy of democratizing information. The amount of harm (direct and indirect) this will cause, and the cost to return to baseline will be so immense, and yet we will not be able to point to the root cause. They won't be there to take responsibility.
Great idea!
Telemarketers have ruined the phone network for me. I haven't answered an unknown call for the past 10 years, which sometimes means I miss important ones. 99.9% of all calls are an attempt to get money, and the 0.1% that's a dentist appointment, a friend that changed numbers or whatever become collateral damage.
A ban is the right idea but I wonder how they can handle it, logistically. I think there needs to be a technical solution.
A national "whitelist", where hospitals, doctors, utility companies and such can register to get their numbers whitelisted perhaps, combined with a setting on phones that block any non-whitelist number.
Each country could maintain their own whitelists, and corrupt nations selling whitelist status to scammers would get blocked in any other country at least.
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.
> These NGOs have converged upon a unified strategy: use the rhetoric of ‘child safety’ to advocate for digital ID laws that would prevent adults from using the internet anonymously.
Of course.
Anyone who brings up kids is trying to manipulate you into giving up your freedom for security. Whatever argument they make should be simply ignored and dismissed.
All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
I'd like to know a lot more about how that works.
A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.
I guess this may be covered by this:
> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;
Linux distro founder here (stagex)
I will never be compelled to implement this, and would never merge it.
Every release requires quorum signatures by an international maintainer team, and the distro is designed to work offline-first, with some variants not even supporting network drivers in the kernel, so Illinois legislators can eat shit.
I feel like all of these laws are being designed backwards. Content providers, like MPAA films, should have to identify what sort of content they are providing. Then I can give my kids a device configured to allow some or all of that at my discretion.
Requiring my kids' devices to advertise their age (or their age "bucket", as if that was a meaningful difference) to protect them is not doing me or my kids any favors.
Disclaimer, I work on Gemma and open models at Deepmind and the opinions here are my own
There were open models from EleutherAI (GPT-Neo), Google Brain (T5X, Bert), and HuggingFace was promoting open models (and others doing open work I haven't listed here) all prior to 2023 and the big Chatgpt moment.
https://github.com/EleutherAI/gpt-neo/releases
https://github.com/google-research/bert
https://github.com/google-research/t5x
If you're learning about AI models it's still worthwhile to review these models and codebases because they continue to be the basis of the technology that's being produced today! It'll give you a good perspective of how things have changed, similar to say learning about propeller planes before moving onto modern jet engines.
That's simply not true. The reason why llama is open source is simply because it got leaked, then llama.cpp was the real game changer which was built from the ground up in depressingly short amount of time. Meta had no choice but to take the L and "support" the open source community. The angry "I-hate-you-and-I-hope-you-die" kind of support.
I think something that doesn't get said enough is Meta did, albeit intentionally kick off the origin of the open source race back in 2023 with the release of llama.
I'm not a big fan of meta in general, but they've done enough good, and it's possible that it was intentional as well. I don't know, I wasn't in the rooms, and I think it's worth giving them some reasonable doubt.
No one is purely good, and no one is purely evil. This is net good regardless.
> Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”). This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
I remain delighted at how absurd our current timeline has become.
You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?
Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."
The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.
Having my name on a bunch of software patents - and, yes, I tried to get my name off them, but was not allowed - I can fairly confidently say: There is not A single worthy software patent out there. You know, one that is "not obvious to someone skilled in the art" and that actually protects a monetary investment.
Software patent are a scourge of the software industry. Patents are designed to protect costly research; simply having an idea is not costly (but it makes in medical research for example). All that software patents do is creating a minefield that hinders competition.
For software Copyright is a far better instrument. Let the one with best implementation win... That's where the cost is: Implementing, testing, shipping, maintaining. Protect that.
Sorry for the rant.
Edit: Spelling
Shrinkflation is all over. Burger buns at fast food places were more dense 20 years ago. The small burgers had 1/8th pound patties instead of the current 1/10th pound. It's hard to find a product that hasn't gotten worse or more expensive, even adjusted for inflation.
My favorite paragraph from Zuckerberg's writeup:
""" [...] it is surprising that the discourse from many developing AI is so filled with doom. I do not understand why anyone who believes that AI will eliminate most jobs and much of humanity's relevance would rush to build that future. The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes. """
Most probably believe this is a good thing, but don't want to give Zuckerberg credit because a) he's had a profoundly negative impact on society and b) the strategy is transparent, he's trying to commoditize his closed rivals, it's not out of principle.
I personally think more open models are a good thing regardless of motive.
Comments here are surprising to me.
I get folks don’t like Zuckerberg and his company and don’t trust his intentions… I don’t either.
But this is an unquestionably good thing right?. The more open source software out there the better. And the more open weights or even over source AI stuff the better too right? More competition the better generally speaking I think.
Unless I’m missing something and am getting this whole situation wrong. Please let me know if I am.
Remember when we needed 200 servers for an enterprise website because Apache used one process or thread per connection - and Nginx collapsed that into a single box overnight? That moment for LLMs is near. It’s going to move us from the big iron era of AI to small portable brains. Nature has already proved it’s possible with 20 watts and very little heat generation. And I think the data center buildout will end in carnage.
It's hilarious how these companies handle security breaches.
I once reported superadmin user/pass committed to github at a major YC backed background check company I worked at and everyone tried to make it seem like it was my fault.
I had just started working there and found it in the first week.
Anyway, had to show that it was committed by their main Staff engineer 2 years before I even worked there. For 2 years everyone's background check data in the United States that went through this thing - millions per year - thousands of Uber drivers, DoorDash, etc. all were viewable with no clearance. Anyone including overseas contractors, new hires, etc. could just login and check anyone's criminal history.
Reporting it was a disaster. They all tried to cover their asses, this huge drama and hand waving started. They tried to blame anyone and everyone. Eventually it was just AWS fault somehow (it wasn't, the Staff engineer was a dumbass, he committed it to a ruby seed file).
-----
I digress, the CTO didn't respond because he was more worried about how it would make him look. This industry is dead - the wrong people work in it.
It’s rather amusing to me to read comments like this, and then simultaneously whenever a Chinese company or team releases open-weight models or whatever there is a giant round of applause, America is so behind, and there’s nothing but positive things to say about the intelligent, creative, and well-intentioned Chinese engineers (which is true, America certainly doesn’t have a monopoly on great people). Don’t you know? Only China can release good, open weight models and American companies can’t compete. Oh by the way all the spend is for nothing because China alone can release open-weight models thus destroying American AI.
When an American company does anything? Doom. And. Gloom. The engineers? Taken to the slaughterhouse! America? Behind! The public? Bamboozeled!
> This is open weights because Meta couldn't monetize it in any other way than to cloud developer's judgement of their reputation.
I’ve been told over and over this doesn’t matter. Just needs to be cheap and open. Or maybe that’s only when Chyna is involved?
Sorry this post is a bit snarky but it really is something to behold. And certainly I don’t know the OP’s opinions on Chinese open weight models. Perhaps they agree with me.
I lament the comments saying this in any way redeems Meta (the company).
The researchers releasing this stuff have almost nothing to do with Meta other than being bankrolled by the slaughterhouse.
You aren't the customer, you are the pawn in big tech's game of thrones. Your good will is a commodity to be traded, almost literally. It will be used against you the moment it's convenient. This is open weights because Meta couldn't monetize it in any other way than to cloud developer's judgement of their reputation.
But I guess most people just don't care.
I'm glad it's open. It does not make me think any better of Meta.
Will be interesting to see how Qwen3.8 27B compares against this once it releases this week. Seems like dense 30B is back in fashion?
EDIT: An open weight version of Muse Spark 1.2 is going to be released as well:
https://x.com/alexandr_wang/status/2086756152034066792
https://xcancel.com/alexandr_wang/status/2086756152034066792
The comment insinuates that the study did not control for age. The study controlled for age.
It's not terrible to not make a good point, but it's not a good point. The not good points should not be ranked highly.